Repository navigation
Judge PR evals on cost and prompt cache changes - #618
Conversation
|
Review: 1 finding. Posted 1 actionable inline comment. |
Preview:
|
84f4cf1 to
6fc6f33
Compare
|
Review: 1 finding. The existing inline finding about cache-break rates depending on run length still applies to this revision. No additional actionable issues found. |
6fc6f33 to
8e1374a
Compare
|
Review: 1 finding. The existing cache-break finding still applies: additional fully cached steps lower the reported rate without improving cache behavior, despite the new instructions claiming independence from run length. No additional actionable issues found. |
The cache hit rate also moves with how long a run is and how much new content it reads, so caching fixes were lost in its noise. comparison.json now carries a cache break rate (tokens the previous step sent that a step sent again instead of reading from cache) and per-run Mann-Whitney p-values for cache hits and breaks. The table bolds a significant cache hit change, and Bonk counts cost and cache changes in its verdict.
8e1374a to
1e74897
Compare
|
LGTM! |
Eval resultsVerdict: ⚪ Unchanged. Nothing the evals run changed, so every result is reused.
|
🔬 Eval runs reviewPerformanceAll four tasks reuse the same results on main and the PR, each passing 10/10 runs; trajectories, costs, durations and tool errors are identical. Cache hits remain 83% for change-calendar, 97% for chess, 93% for incident-desk and 90% for worker-logs, with unchanged cache-break rates and no comparable independent runs demonstrating a caching gain. Chess is the most expensive task at 66.9 model steps and 11.2 minutes per run, costing $0.0483–$0.0854; implementation and testing dominate its steps, with malformed calls and exact-text edit retries adding avoidable work. ⚪ VERDICT: NO REGRESSION FROM THIS PRThe diff changes comparison reporting and advisory-review instructions, not what the agents execute, so the reused results support neither an agent-performance regression nor an improvement. TriageTool errors
What to do
|
* Describe Google Chat bindings by what they are bound to (cloudflare#606) An agent holding a Chat Conversation binding (a ChatSpace) took it for an account session and called searchSpaces() on it before recovering. The observation opening every Chat binding was titled "Open a Google Chat session", which reinforced the guess. - ChatSession, ChatSpace and ChatThread now say which resource each session is bound to, as the Slack workspace, conversation and thread sessions do, and ChatSpace and ChatThread point at post(). - The opening observation names what it opened: a Chat session, conversation or thread. * Review with GPT-6.1 Sol and opencode 1.18.33 in Bonk (cloudflare#607) Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Say a new Gadget starts empty, and name missing files in editFile (cloudflare#600) In 24 of 40 eval trials, the agent read server.js and client.js right after creating an empty Gadget, because the system prompt said every Gadget has both files. The prompt now says to create them, and that a new Gadget has none unless it came from a blueprint. editFile checked that the agent had read a file before checking that the file exists, so a mistyped name got "You must read a file before you can edit it." It now names the missing file. The check reuses readFile's lookup, moved into readToolFile, and runs only when the edit is refused. The editFile description now states the read rule. * Let Bonk judge whether a PR improves the eval runs (cloudflare#608) * Let Bonk judge whether a PR improves the eval runs Bonk's verdict followed the pass-rate test alone, so a PR that only made the agent's work better came back as "no regression". cloudflare#600 cut the runs that read files in a new, empty Gadget from 20 of 40 to 2 of 40, and cloudflare#601 raised cache hits on every task; both got a ⚪. Bonk now decides from both sides' trajectories as well as comparison.json: pass rates, wasted steps, tool errors, cache hit rate, time and what the agent built. A change, better or worse, counts only when the diff explains it, both sides show it, and it is larger than run-to-run variation. The Tool errors and Prompt cache sections now show both sides, and the cache section reports rises as well as falls. * Call a mixed eval outcome inconclusive, not a regression The regression rule came first and fired on any change for the worse, so a PR that also counted a change for the better could never reach the inconclusive case meant for it. * Read the Bonk model from the BONK_MODEL repository variable The three Bonk workflows each named their model twice, and the eval review had fallen behind on gpt-5.6-sol. BONK_MODEL is set to openai/gpt-6.1-sol. * Bump dompurify from 3.4.15 to 3.4.16 (cloudflare#614) Bumps [dompurify](https://github.com/cure53/DOMPurify) from 3.4.15 to 3.4.16. - [Release notes](https://github.com/cure53/DOMPurify/releases) - [Commits](cure53/DOMPurify@3.4.15...3.4.16) --- updated-dependencies: - dependency-name: dompurify dependency-version: 3.4.16 dependency-type: direct:development ... Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> * Suggest GPT-6.1 Sol and Claude Sonnet 5.5 and bump pi to 0.99.1 (cloudflare#610) * Upgrade pi to 0.99.1 * Add GPT-6.1 Sol and Claude Sonnet 5.5 to the suggested models * Deflake the scheduler callback concurrency test (cloudflare#617) The test required four 10ms callbacks to overlap by chance; on a loaded runner the startHook and authorization round trips could outlast that window. Hold callbacks until the concurrency bound has been reached once. Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Resume agent turns after auto-approval, and fix retryAgent and Scheduler docs (cloudflare#599) * Document that retryAgent throws while an agent is running The implementation rejects with the agent-running error, as the other chat mutators do; the doc said it silently did nothing. * Show executeCode minting scheduler callbacks via env.GADGET[restore] The Scheduler's agent-facing types told agents to call ctx.restore() from executeCode, which targets the executeCode worker and throws that it implements no [restore](). The example also called a bare scheduler, the Workers global, rather than env.SCHEDULER. * Resume agent turns whose awaited actions an auto-approval drain decides Only a manual approveAction resumed a turn suspended on awaitDecision. "Always approve" (the chat card, Activity panel and Connections rule toggle) and approveAction's cascade apply actions through the drain alone, so when the drain decided a turn's last awaited action -- e.g. the first Gmail archive or a vetted MCP tool call -- the action applied and the agent never resumed or learned of it. Both drains now resume the chats whose awaited actions were pending on that gatekeeper. A drain requested while one is running now shares its promise, so awaiting drain() waits for the rerun it requested instead of returning before the action is applied. No caller awaits a drain, so a failed resume is logged rather than ending the loop. An auto-approvable awaited action queued behind a manual gate, which the in-order drain stops at, also let its turn run on while the action sat unapplied. submitAction now suspends that turn too, and rejectAction drains as approveAction does, so rejecting the gate applies the queued action and resumes the turn instead of leaving both waiting on a drain that nothing starts. * Resume only chats whose awaited actions the drain approved The drain-and-resume path checked every chat with an awaited action pending on the gatekeeper, including ones the drain left pending, so a chat whose older action stayed on a manual gate could be resumed from its latest turn's already-approved actions. * Migrate Worker configs to cloudflare.config.ts (cloudflare#597) * Author Worker configs in cloudflare.config.ts Each Worker's config is now a TypeScript cloudflare.config.ts in the @cloudflare/config format, built from a shared factory in scripts/worker-config.ts that holds the compatibility date, the observability block, the capnweb-validate build and the Text module rules the configs used to copy. Wrangler's --experimental-new-config covers only dev, build, deploy and versions, and rejects --config, so it can't drive the multi-worker dev server, wrangler types, the vitest pools or the integration harness. The wrangler.jsonc beside each config therefore stays, committed but generated by scripts/generate-worker-configs.ts, and every existing consumer reads it unchanged. Settings the new format has no field for (build, rules, KV preview_id, assets.directory) go in a named `wrangler` export, and Durable Object migrations stay tagged in a `migrations` export, because the release manifest ships them to the deploy service verbatim. The generated files parse identical to the hand-written ones minus $schema, and the golden manifest is unchanged. `pnpm configs:check` fails on a stale file, a hand-written wrangler.jsonc with no source, or an export the generator doesn't read; it runs in `pnpm lint` and CI's lint job, and `pnpm dev-server` regenerates before starting. * Share the Workers compatibility date with the vitest pools The Worker vitest pools copied the deployed compatibility date by hand. They now import COMPATIBILITY_DATE from @gadgets/scripts/worker-config, so a date bump reaches the runtime the tests model too. Flags stay per-pool, since several pools deliberately differ from the deployed Workers. The library pools (observability, gatekeeper-kit, typed-storage) keep their literal: they model no deployed Worker. * Point docs and the gatekeeper skill at cloudflare.config.ts The write-gatekeeper skeleton now shows a cloudflare.config.ts built from the shared factory, and the skill registers a new gatekeeper in workshop-backend's env with bindings.worker(). AGENTS.md describes the generated wrangler.jsonc and the configs:generate/configs:check pair, and the READMEs that pointed at wrangler.jsonc for a flag or var now point at its source. * Trim worker-config duplication left by the migration - Drop dev-router's enable_ctx_exports flag, the default since 2025-11-17 (wrangler dev warned it was redundant). - Add DEFAULT_GATEKEEPER_WRANGLER for the capnweb-validate build plus .txt/.svg Text rules that 13 gatekeepers spelled out; their generated wrangler.jsonc is unchanged. - The six vitest pools that model their deployed runtime exactly read compatibilityDate and compatibilityFlags from their cloudflare.config.ts instead of copying them. The router and backend pools, whose flags differ on purpose, keep theirs. * Hide superseded models from pickers while they still resolve (cloudflare#611) * Hide superseded models from pickers while they still resolve SUGGESTED_MODELS is also the AI Gateway registry, so stored model ids (chat authors, spawner configs, preferences) must keep resolving. A hidden model is left out of getModelList() but resolveModel() still finds it, and an external message to a chat on a hidden model stays on it instead of falling back to the preferred model. * Keep dialog actions within the viewport (cloudflare#619) The "Verify your access" dialog lists one card per connection, and Kumo's Dialog sets no max-height, so with enough connections its "Verify and open" button was pushed below the fold on desktop and only zooming out reached it. The Cloudflare account picker (one card per account) and the Blueprint settings dialog (a textarea) could overflow the same way. All three now use the bounded layout BlueprintModal adopted in cloudflare#436: a capped flex-column dialog with a fixed header and footer and a scrolling body. The mobile bottom-sheet rule for responsive-dialog still overrides the top offset and max-height on phones. * Public-API kernel integration tests (round 5) (cloudflare#620) * Stop proposing a removed gadget in chats that pinned it * Add round-5 kernel e2e scenarios * Harden round-5 kernel tests and share fixture control helpers * Assert only intended behaviour; keep the concurrent-approval race as an it.fails repro * Bump @cloudflare/workers-types (cloudflare#622) Bumps the cloudflare-toolchain group with 1 update in the / directory: [@cloudflare/workers-types](https://github.com/cloudflare/workerd). Updates `@cloudflare/workers-types` from 5.20260903.1 to 5.20260924.1 - [Release notes](https://github.com/cloudflare/workerd/releases) - [Changelog](https://github.com/cloudflare/workerd/blob/main/RELEASE.md) - [Commits](https://github.com/cloudflare/workerd/commits) --- updated-dependencies: - dependency-name: "@cloudflare/workers-types" dependency-version: 5.20260924.1 dependency-type: direct:production update-type: version-update:semver-minor dependency-group: cloudflare-toolchain ... Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> * Bump the github-actions group with 3 updates (cloudflare#621) Bumps the github-actions group with 3 updates: [voidzero-dev/setup-vp](https://github.com/voidzero-dev/setup-vp), [actions/upload-artifact](https://github.com/actions/upload-artifact) and [actions/download-artifact](https://github.com/actions/download-artifact). Updates `voidzero-dev/setup-vp` from 1.17.0 to 1.21.1 - [Release notes](https://github.com/voidzero-dev/setup-vp/releases) - [Commits](voidzero-dev/setup-vp@v1.17.0...3754dd7) Updates `actions/upload-artifact` from 4.6.2 to 7.0.1 - [Release notes](https://github.com/actions/upload-artifact/releases) - [Commits](actions/upload-artifact@v4.6.2...043fb46) Updates `actions/download-artifact` from 5.0.0 to 8.0.1 - [Release notes](https://github.com/actions/download-artifact/releases) - [Commits](actions/download-artifact@634f93c...3e5f45b) --- updated-dependencies: - dependency-name: voidzero-dev/setup-vp dependency-version: 1.21.1 dependency-type: direct:production update-type: version-update:semver-minor dependency-group: github-actions - dependency-name: actions/upload-artifact dependency-version: 7.0.1 dependency-type: direct:production update-type: version-update:semver-major dependency-group: github-actions - dependency-name: actions/download-artifact dependency-version: 8.0.1 dependency-type: direct:production update-type: version-update:semver-major dependency-group: github-actions ... Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> * Bump yjs from 13.6.32 to 13.6.33 in the collab-yjs group (cloudflare#626) Bumps the collab-yjs group with 1 update: [yjs](https://github.com/yjs/yjs). Updates `yjs` from 13.6.32 to 13.6.33 - [Release notes](https://github.com/yjs/yjs/releases) - [Commits](yjs/yjs@v13.6.32...v13.6.33) --- updated-dependencies: - dependency-name: yjs dependency-version: 13.6.33 dependency-type: direct:production update-type: version-update:semver-patch dependency-group: collab-yjs ... Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> * Bump the react-and-ui group with 3 updates (cloudflare#624) Bumps the react-and-ui group with 3 updates: [@cloudflare/kumo](https://github.com/cloudflare/kumo/tree/HEAD/packages/kumo), [@tanstack/react-router](https://github.com/TanStack/router/tree/HEAD/packages/react-router) and [@tanstack/router-plugin](https://github.com/TanStack/router/tree/HEAD/packages/router-plugin). Updates `@cloudflare/kumo` from 2.13.2 to 2.14.0 - [Release notes](https://github.com/cloudflare/kumo/releases) - [Changelog](https://github.com/cloudflare/kumo/blob/main/packages/kumo/CHANGELOG.md) - [Commits](https://github.com/cloudflare/kumo/commits/@cloudflare/kumo@2.14.0/packages/kumo) Updates `@tanstack/react-router` from 1.170.33 to 1.170.39 - [Release notes](https://github.com/TanStack/router/releases) - [Changelog](https://github.com/TanStack/router/blob/main/packages/react-router/CHANGELOG.md) - [Commits](https://github.com/TanStack/router/commits/@tanstack/react-router@1.170.39/packages/react-router) Updates `@tanstack/router-plugin` from 1.168.36 to 1.168.40 - [Release notes](https://github.com/TanStack/router/releases) - [Changelog](https://github.com/TanStack/router/blob/main/packages/router-plugin/CHANGELOG.md) - [Commits](https://github.com/TanStack/router/commits/@tanstack/router-plugin@1.168.40/packages/router-plugin) --- updated-dependencies: - dependency-name: "@cloudflare/kumo" dependency-version: 2.14.0 dependency-type: direct:production update-type: version-update:semver-minor dependency-group: react-and-ui - dependency-name: "@tanstack/react-router" dependency-version: 1.170.39 dependency-type: direct:production update-type: version-update:semver-patch dependency-group: react-and-ui - dependency-name: "@tanstack/router-plugin" dependency-version: 1.168.40 dependency-type: direct:development update-type: version-update:semver-patch dependency-group: react-and-ui ... Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> * Bump the editor-codemirror group across 1 directory with 4 updates (cloudflare#625) Bumps the editor-codemirror group with 4 updates in the / directory: [@codemirror/commands](https://github.com/codemirror/commands), [@codemirror/state](https://github.com/codemirror/state), [@codemirror/view](https://github.com/codemirror/view) and [@lezer/highlight](https://github.com/lezer-parser/highlight). Updates `@codemirror/commands` from 6.11.0 to 6.11.1 - [Changelog](https://github.com/codemirror/commands/blob/main/CHANGELOG.md) - [Commits](https://github.com/codemirror/commands/commits) Updates `@codemirror/state` from 6.7.4 to 6.7.6 - [Changelog](https://github.com/codemirror/state/blob/main/CHANGELOG.md) - [Commits](https://github.com/codemirror/state/commits) Updates `@codemirror/view` from 6.43.11 to 6.43.13 - [Changelog](https://github.com/codemirror/view/blob/main/CHANGELOG.md) - [Commits](https://github.com/codemirror/view/commits) Updates `@lezer/highlight` from 1.2.3 to 1.2.4 - [Changelog](https://github.com/lezer-parser/highlight/blob/main/CHANGELOG.md) - [Commits](https://github.com/lezer-parser/highlight/commits) --- updated-dependencies: - dependency-name: "@codemirror/commands" dependency-version: 6.11.1 dependency-type: direct:production update-type: version-update:semver-patch dependency-group: editor-codemirror - dependency-name: "@codemirror/state" dependency-version: 6.7.6 dependency-type: direct:production update-type: version-update:semver-patch dependency-group: editor-codemirror - dependency-name: "@codemirror/view" dependency-version: 6.43.13 dependency-type: direct:production update-type: version-update:semver-patch dependency-group: editor-codemirror - dependency-name: "@lezer/highlight" dependency-version: 1.2.4 dependency-type: direct:production update-type: version-update:semver-patch dependency-group: editor-codemirror ... Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> * Bump the remaining-npm group with 7 updates (cloudflare#627) * Bump the remaining-npm group with 7 updates Bumps the remaining-npm group with 7 updates: | Package | From | To | | --- | --- | --- | | [isomorphic-git](https://github.com/isomorphic-git/isomorphic-git) | `1.41.9` | `1.42.2` | | [yaml](https://github.com/eemeli/yaml) | `2.9.0` | `2.9.1` | | [motion](https://github.com/motiondivision/motion) | `13.2.0` | `13.4.2` | | [diff3](https://github.com/axosoft/diff3) | `0.0.3` | `0.0.4` | | [vitest-evals](https://github.com/getsentry/vitest-evals/tree/HEAD/packages/vitest-evals) | `0.16.1` | `0.17.0` | | [@oxlint/plugins](https://github.com/oxc-project/oxc/tree/HEAD/npm/oxlint-plugins) | `1.73.0` | `1.85.0` | | [zod](https://github.com/colinhacks/zod) | `4.5.4` | `4.6.5` | Updates `isomorphic-git` from 1.41.9 to 1.42.2 - [Release notes](https://github.com/isomorphic-git/isomorphic-git/releases) - [Commits](isomorphic-git/isomorphic-git@v1.41.9...v1.42.2) Updates `yaml` from 2.9.0 to 2.9.1 - [Release notes](https://github.com/eemeli/yaml/releases) - [Commits](eemeli/yaml@v2.9.0...v2.9.1) Updates `motion` from 13.2.0 to 13.4.2 - [Changelog](https://github.com/motiondivision/motion/blob/main/CHANGELOG.md) - [Commits](motiondivision/motion@v13.2.0...v13.4.2) Updates `diff3` from 0.0.3 to 0.0.4 - [Commits](https://github.com/axosoft/diff3/commits) Updates `vitest-evals` from 0.16.1 to 0.17.0 - [Release notes](https://github.com/getsentry/vitest-evals/releases) - [Commits](https://github.com/getsentry/vitest-evals/commits/v0.17.0/packages/vitest-evals) Updates `@oxlint/plugins` from 1.73.0 to 1.85.0 - [Release notes](https://github.com/oxc-project/oxc/releases) - [Changelog](https://github.com/oxc-project/oxc/blob/main/CHANGELOG.md) - [Commits](https://github.com/oxc-project/oxc/commits/oxlint_v1.85.0/npm/oxlint-plugins) Updates `zod` from 4.5.4 to 4.6.5 - [Release notes](https://github.com/colinhacks/zod/releases) - [Commits](colinhacks/zod@v4.5.4...v4.6.5) --- updated-dependencies: - dependency-name: isomorphic-git dependency-version: 1.42.2 dependency-type: direct:production update-type: version-update:semver-minor dependency-group: remaining-npm - dependency-name: yaml dependency-version: 2.9.1 dependency-type: direct:production update-type: version-update:semver-patch dependency-group: remaining-npm - dependency-name: motion dependency-version: 13.4.2 dependency-type: direct:production update-type: version-update:semver-minor dependency-group: remaining-npm - dependency-name: diff3 dependency-version: 0.0.4 dependency-type: direct:production update-type: version-update:semver-patch dependency-group: remaining-npm - dependency-name: vitest-evals dependency-version: 0.17.0 dependency-type: direct:production update-type: version-update:semver-minor dependency-group: remaining-npm - dependency-name: "@oxlint/plugins" dependency-version: 1.85.0 dependency-type: direct:development update-type: version-update:semver-minor dependency-group: remaining-npm - dependency-name: zod dependency-version: 4.6.5 dependency-type: direct:production update-type: version-update:semver-minor dependency-group: remaining-npm ... Signed-off-by: dependabot[bot] <support@github.com> * Keep diff3 at 0.0.3 and @oxlint/plugins at vite-plus's pin diff3 0.0.4 has a known issue, so revert to 0.0.3 and have Dependabot skip that version (later releases are still offered). @oxlint/plugins is types-only and must equal the version vite-plus pins (scripts/oxlint-plugin.test.ts); 1.85.0 failed that check against vite-plus 0.2.8's =1.73.0. Revert it and leave it to move by hand with vite-plus, since Dependabot always picks the newest release rather than the pinned one. --------- Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> Co-authored-by: Nathan Disidore <nathan@cloudflare.com> * Bump node-html-parser from 7.1.0 to 9.0.4 (cloudflare#629) Bumps [node-html-parser](https://github.com/taoqf/node-fast-html-parser) from 7.1.0 to 9.0.4. - [Release notes](https://github.com/taoqf/node-fast-html-parser/releases) - [Changelog](https://github.com/taoqf/node-html-parser/blob/main/CHANGELOG.md) - [Commits](taoqf/node-html-parser@v7.1.0...v9.0.4) --- updated-dependencies: - dependency-name: node-html-parser dependency-version: 9.0.4 dependency-type: direct:production update-type: version-update:semver-major ... Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> * Bump pako and @types/pako (cloudflare#630) Bumps [pako](https://github.com/nodeca/pako) and [@types/pako](https://github.com/DefinitelyTyped/DefinitelyTyped/tree/HEAD/types/pako). These dependencies needed to be updated together. Updates `pako` from 2.2.0 to 3.0.2 - [Changelog](https://github.com/nodeca/pako/blob/master/CHANGELOG.md) - [Commits](nodeca/pako@2.2.0...3.0.2) Updates `@types/pako` from 2.0.4 to 3.0.0 - [Release notes](https://github.com/DefinitelyTyped/DefinitelyTyped/releases) - [Commits](https://github.com/DefinitelyTyped/DefinitelyTyped/commits/HEAD/types/pako) --- updated-dependencies: - dependency-name: pako dependency-version: 3.0.2 dependency-type: direct:production update-type: version-update:semver-major - dependency-name: "@types/pako" dependency-version: 3.0.0 dependency-type: direct:development update-type: version-update:semver-major ... Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> * Explain why the code view is read-only outside a conversation (cloudflare#633) Code edits are made on a chat's branch, so the Code view locks editing when no conversation is selected -- e.g. a workspace freshly created from a blueprint, or one reopened from the sidebar. Nothing said so: the header just read "Viewing" and the New file button was silently disabled, which users read as the editor breaking until they sent a new prompt. Show "Select or start a conversation to edit" as a banner above the editor and as the disabled New file button's tooltip. * Bump vite-plus to 1.0.0 (cloudflare#632) vite-plus 1.0.0 is the only VoidZero package that can move within the 24h minimumReleaseAge window without breaking the toolchain: - vite stays 7.3.6: Vite 8's Oxc leaves Stage-3 decorators unlowered, so every workerd suite importing a `@validateRpc()` class fails with "SyntaxError: Invalid or unexpected token" (re-verified on 8.3.1; see the catalog comment). - vitest stays ^4.1.11: @cloudflare/vitest-pool-workers 0.22.0 peers vitest ^4.1.0. - @vitejs/plugin-react stays ^5.2.0: 6.x requires vite ^8. Migration for 1.0: - Override `vite@*` instead of bare `vite`, so vite-plus keeps its `vite` -> @voidzero-dev/vite-plus-core alias instead of having vite 7 substituted under `vp`. - Move task `env`/`input`/`output` under `cache` (vite-task#749), in the package configs and the shared task builders. - Import the oxlint plugin types and RuleTester from `vite-plus/lint/plugins{,-dev}` and drop the separately pinned @oxlint/plugins dependency, its drift test and dependabot ignore. - oxlint 1.85: use toSorted() on three freshly built arrays, and turn off react/globals for workshop-frontend tests, whose hook probes assign to an outer `let` on purpose. * Bump the build-toolchain group across 1 directory with 2 updates (cloudflare#635) Bumps the build-toolchain group with 2 updates in the / directory: [esbuild](https://github.com/evanw/esbuild) and [terser](https://github.com/terser/terser). Updates `esbuild` from 0.28.1 to 0.28.2 - [Release notes](https://github.com/evanw/esbuild/releases) - [Changelog](https://github.com/evanw/esbuild/blob/main/CHANGELOG.md) - [Commits](evanw/esbuild@v0.28.1...v0.28.2) Updates `terser` from 5.49.2 to 5.51.2 - [Changelog](https://github.com/terser/terser/blob/master/CHANGELOG.md) - [Commits](terser/terser@v5.49.2...v5.51.2) --- updated-dependencies: - dependency-name: esbuild dependency-version: 0.28.2 dependency-type: direct:development update-type: version-update:semver-patch dependency-group: build-toolchain - dependency-name: terser dependency-version: 5.51.2 dependency-type: direct:development update-type: version-update:semver-minor dependency-group: build-toolchain ... Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> * Refactor / cleanup: Pull storage schemas out into their own files. (cloudflare#639) This is the first, small step to break up overseer.ts, and probably the most obvious. All of the typed-storage collections and TS types are now defined under a `storage-schema` subdirectory. I generally think of APIs and storage schemas as defining the architecture of a piece of software. Everything in between is just glue, and can change more easily. Consolidating storage schemas into one place makes it easy to tell what storage schemas are changing in any given PR. I also moved migrations into this new directory. Claude suggested several small related cleanups as well which I told it to go ahead and do. * Add the eval user model only for direct access (cloudflare#634) Since cloudflare#611, addModel refuses an id a gateway model would shadow, and the eval target added one in gateway mode, where the gateway already serves the eval model. So every eval trial failed at setup. * Gmail gatekeeper: Fully simulate actions (cloudflare#638) * Plan for improving gmail simulation. * Make gmail gatekeeper simulate label changes. This is part 1 of plans/gmail-simulation.md. * Make gmail gatekeeper simulate sends. This is part 2 of plans/gmail-simulation.md. * Judge PR evals on cost and prompt cache changes (cloudflare#618) The cache hit rate also moves with how long a run is and how much new content it reads, so caching fixes were lost in its noise. comparison.json now carries a cache break rate (tokens the previous step sent that a step sent again instead of reading from cache) and per-run Mann-Whitney p-values for cache hits and breaks. The table bolds a significant cache hit change, and Bonk counts cost and cache changes in its verdict. * Send the system prompt as a static and a dynamic block (cloudflare#609) pi sends the leading system message as one block, so a change to the project-specific text missed the cache from the start of the prompt. The handle now splits it after the static text, with a cache breakpoint there: on Anthropic by moving the system block's breakpoint, and on OpenAI GPT-5.6+ with an explicit prompt_cache_breakpoint. * Upload preview secrets to the Preview base config (cloudflare#648) The Workers API stopped accepting `preview_defaults`, the field the pinned draft Wrangler writes for `wrangler preview secret bulk`, so every preview deploy failed at the first worker with secrets. The secrets now go through `wrangler preview base-config secret bulk` on the workspace's own Wrangler, which writes `previews_base_config`. Everything else stays on the draft build: neither sibling preview bindings nor per-preview KV and R2 provisioning is in a released Wrangler, checked against 4.138.0 and 4.147.0. * Public-API kernel integration tests (round 6) (cloudflare#647) Cover user journeys a kernel refactor could silently break: - the agent's system prompt follows the workspace's gadgets, bindings, ambient connection and standard formats, and never another chat's pending binding - a pasted link (capsule) becomes a binding for that chat only - a chat attachment reaches the model and stays in later turns - one accept commits every gadget a chat built, each with its own code - gadget console logs reach subscribers labelled draft or mainline, stop on dispose, and never reach a use collaborator - a gadget's LLM binding runs on its bound model; a blueprint install uses the installer's model - admin instance instructions and format hints reach the agent, not the user - disconnecting an account from the Connectors page removes and revokes it Adds systemPromptOf() to the mock model for reading recorded prompts. Also makes bare `.rejects.toThrow()` assertions on RPC promises able to fail. vitest's `.rejects` calls a callable subject, and a Cap'n Web RpcPromise is callable: calling it pipelines a call on the result, which rejects with "'' is not a function." when the original call succeeded, so those assertions passed whatever the call did. They now match the refusal message; two were also wrong, since writeValue resolves once the write is submitted for approval, and now assert that instead. * Restart a workspace whose loop counter is exhausted (cloudflare#640) * Restart a workspace whose loop counter is exhausted The Workers runtime refuses a Durable Object call once the loop counter behind it is spent ("Subrequest depth limit exceeded. This request looped back into the Workers runtime too many times."). A workspace object's outgoing channels can end up holding a spent counter with nothing recursing, and from then on every call it makes to a user object is refused until the instance is replaced. When one of the workspace's own user-object calls is rejected that way (a call through a wrapped stub, the last-active bump, or the outputs sync), the workspace now schedules the existing access restart. At most one restart per instance, and none in an instance's first 60 seconds. Errors thrown by gadget, agent or gatekeeper-facet code are never consulted. * Let admins manage a deployment's AI Gateway models via admin panel (cloudflare#616) A deployment could only change which models its AI Gateway offers by patching SUGGESTED_MODELS. That patch goes stale whenever the catalog changes, and some models can't go in the public catalog at all. In AI Gateway mode, /admin gets a new Models tab: Enable / test providers and set default reasoning for the deployment and more! * feat(google-gatekeeper): begin Google Chat DMs and group chats from the account binding (cloudflare#649) * Start Google Chat direct messages and group chats from the account binding The whole-account Chat session gains two methods: - searchPeople(query) searches the connected account's Workspace directory (domain profiles only, never contacts), recording each page as an observation. - sendDirectMessage(people, text) sends to one person (a DM) or 2-49 people (a group chat with exactly them). An existing conversation is an ordinary send. Otherwise the message is queued as a new "chatStartConversation" action kind, separately auto-approvable from sends, and every person must be named by email and resolve to a directory profile: outsiders and yourself are refused before anything is queued. Applying a start creates the conversation idempotently (spaces.setup with a requestId, remembered once created), reuses a group chat that appeared since queueing, and verifies membership before posting, since Google silently drops anyone who blocks the caller from a new group chat. Until committed, the message carries a temporary pending:space:{id}; edits queued against it carry over into the real conversation, replies are refused. Undo deletes the message. The account resource adds chat.spaces.create (not chat.spaces) and directory.readonly, so existing account connections re-consent. * Fix Google Chat sends whose new conversation exists before they post A send that starts a direct message or group chat is queued under a temporary pending:space name. If applying it set up the conversation but the post then failed, two things went wrong: - ChatSpace.listMessages() for the new conversation left the queued send out, so an agent could conclude it was never sent and send it again. listForSpace and resolveMessage now resolve the temporary name to the conversation, which also lets the send's capability pass that conversation's scope check. The send reports the real spaceId once its conversation exists, and the agent-facing docs say so. - A failed member check after a successful setup, such as a 503, left the send marked as possibly sent, so it could not be rejected until a retry succeeded. The setup's mark is now cleared as soon as the conversation exists: an empty conversation shows nobody anything. A retry after the conversation is recorded still skips setup, so a send whose post may have landed stays unrejectable. A new test pins that. * Export gatekeeper-kit's SingleFlight and use it for Chat conversation setup SingleFlight coalesces concurrent work by key and releases each flight once it settles. It was internal to the kit; it is now the ./single-flight subpath, listed in the README inventory and noted in the design record. The Google Chat gatekeeper used a hand-rolled map of promises to make sends to the same people, applied at once, share one conversation setup. It now uses SingleFlight, whose own tests cover joining and release. * Re-check a new Google Chat conversation's members before retrying an unsent post A send that starts a group chat records the conversation once it is set up. If the post was then refused outright, a retry reused that conversation without checking who was in it, so someone who joined in between would receive the approved message. A retry that definitely didn't post now refuses when the members have changed; one whose post may have landed still goes back to the same conversation, so it stays idempotent. The directory lookup that confirms each person in a new conversation read one page of matches and called anyone not on it outside the organization. When Google reports further pages it now says it couldn't confirm them instead. * Check that Google Chat sends reach exactly the approved people A new conversation's send checked only that nobody approved was missing, and several paths could still reach someone who wasn't: - A replayed spaces.setup returns the group as it is now, so after a lost setup response someone added since would receive the message. The send is now refused when the conversation holds anyone extra, and stays rejectable. - A retry after a post that may have landed skipped the member check. Every retry that reuses the recorded conversation now checks it first, leaving the attempt mark as it is, so an uncertain send stays unrejectable. - findGroupChat trusted a matching member count, but Google's fallback for a block can offer a group with someone else in it. Each requested person is now confirmed, by user ID against the member list or by email through members.get, and a membership Google reports as NOT_A_MEMBER doesn't count. - peopleIn read only the first page of members. It now follows every page, since Google may return fewer members than asked for. * Simplify Google Chat conversation starts A send that creates its conversation is queued under a temporary pending:space: name. That name was resolved at five call sites; the store now resolves it once, in get() and list(), so every reader sees the real conversation once it exists. The guard that keeps such a send from being replied to keyed on the temporary spaceId, which missed the window after setup but before the post; it now checks the stored send. openConversation returns as soon as Google's lookup finds the group chat, since the lookup already confirmed exactly who is in it, rather than listing the members a second time. Its post-setup check is one early return and one throw instead of a nested ternary. The directory search reads its response through readGoogleJson, as the other People API call does, so it is size-bounded and logs Google's reasons; the #request/#fetchJson split that existed only for it is gone. The 400/404 "no such user" check shared by three lookups is one predicate, and #onlyIn no longer reads a membership for a users/{id} reference that is already absent. The tests drop unused backend state and share the group-chat setup and Chat-app membership they repeated. * Address Codex P2 review findings on sync PR - VoiceSettingsDialog: clear the saved speaker when switching the conversation TTS model, so validation resolves the new model's default voice instead of rejecting the whole update. - overseer #drainAutoApprovalsAndResume: snapshot non-agent awaited actions too and fan out per approved action as approveAction does, so a drain that auto-applies a user/OpenAPI-caller action resumes the turn waiting on it. - overseer loop-limit recovery: route every User-DO call's rejections through recovery -- wrapUserDo at 14 construction sites (get/getByName) plus explicit restartIfLoopLimited in the catches covering only User-DO calls. Documents the invariant on restartIfLoopLimited. - user getModelReasoning: borrow behavesLike levels for added models the runtime doesn't know, matching gatewayCatalogModel, so the composer shows their effort selector. Tests: 2 dialog cases, drain-resume (negative-controlled), behavesLike borrow/no-borrow/own-wins, 8 loop-limit routes (spot-negative-controlled); openFakeOverseer gains a ctx passthrough for first-open coverage. --------- Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: Nathan Disidore <nathan@cloudflare.com> Co-authored-by: Maximo Guk <62088388+Maximo-Guk@users.noreply.github.com> Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Co-authored-by: Ashish Kumar Singh <ashishsingh@cloudflare.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> Co-authored-by: Kenton Varda <kenton@cloudflare.com>
PR eval comments hide caching wins. #609 removed most cache misses at turn starts, but the comment showed +1 to +3 pp in cache hits, and Bonk called it unchanged.
The cache hit rate is cache reads over all prompt tokens. It also moves with how long a run is and how much new content it reads, so a real caching fix gets lost in run-to-run noise.
Changes
comparison.jsongetscacheBreakRateper task (of the tokens the previous step had already sent, the share a step sent again instead of reading from cache), plus p-values for cache hits and cache breaks. Each is an exact Mann-Whitney U test on each run's own rate, in the direction the pooled rate moved.On #609's results, cache breaks fell on all 3 comparable tasks (p = 0.029, 0.002, 0.075).
comparison.tsonly builds reports, so no evals rerun.